[GLM-5.2 GB200] Preserve draft precision and disable background UCX progress / [GLM-5.2 GB200] 保留 draft 精度并关闭 UCX 后台进展 - #3401
edwingao28 wants to merge 2 commits into
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
91efe8b to
e908cc4
Compare
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36538876933 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36538876933 |
68d3ce2 to
274c69d
Compare
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
6113b30 to
fca8de5
Compare
…refill Keep the draft at shipped precision and disable only the two background UCX progress controls for disaggregated prefill. Preserve images, workload, strict synchronization and golden acceptance. Request the scoped engine-patch waiver. 保留 GLM-5.2 GB200 draft 的原始精度,仅关闭分离式 prefill 的两项 UCX 后台进展控制;保留镜像、工作负载、严格同步及 golden acceptance,并申请对应补丁例外。
fca8de5 to
9fbe9bf
Compare
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
中文:同步 main,并将原有 UCX setup 选择迁移到 Python launcher。[skip-sweep]
|
Claude finished @edwingao28's task in 47s —— View job LGTM, no blocking issues found. I followed the setup-script choice from start to finish.
I couldn't run |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked whether the master config (nvidia-master.yaml) needed a matching update for the SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE flip — it doesn't set that env var for these four recipe keys, so there's no master-config drift to fix.
Extended reasoning...
The diff touches four GLM-5.2 GB200 recipe YAMLs (compliant 1->0 flip of a prohibited draft-precision flag), a new hash-guarded NIXL conn.py patch script plus launcher plumbing (lanes.py/submit.py) to select it for two disaggregated recipes, a pending-approval engine-patch waiver doc, new driver tests, and a perf-changelog append. An inline finding already flags a policy-violating Chinese description line in the new perf-changelog entry, so a human look is warranted regardless; I additionally verified there's no corresponding nvidia-master.yaml entry for the flag that would need updating in lockstep, ruling out that specific cross-file consistency concern.
| - "Keep the GLM-5.2 NextN/MTP draft at shipped precision and disable background UCX progress for GB200 disaggregated prefill to avoid the observed event-arm crash; images, topology, workload and golden acceptance are unchanged." | ||
| - "使 GLM-5.2 NextN/MTP draft 保持原始发布精度,并关闭 GB200 分离式 prefill 的 UCX 后台进展以规避已观察到的 event-arm 崩溃;镜像、拓扑、工作负载和 golden acceptance 不变。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3401 |
There was a problem hiding this comment.
🟡 (optional) This new perf-changelog entry adds a Chinese description line, which AGENTS.md:126 explicitly prohibits for new entries (English-only, no bilingual descriptions), unlike the rest of the bilingual-docs policy. Fix: remove the Chinese description string and keep only the English description line in this appended entry, consistent with every other entry in the file.
Why this was flagged
AGENTS.md:126 states 'New inferencex-e2e/perf-changelog.yaml entries must be English-only. Do not add Chinese translations or bilingual descriptions.' The new entry appended at inferencex-e2e/perf-changelog.yaml:9247-9249 includes both an English description and a Chinese description ('使 GLM-5.2 NextN/MTP draft 保持原始发布精度...') under the same description: list. This is a direct violation of the file's documented English-only invariant, which the rest of the bilingual-docs policy (AGENTS.md:20) does not override since the comment at line 126 explicitly carves this file out. No other check in the diff catches this since the file is otherwise append-only/byte-sensitive and no linter is shown running on it.
Verification: nit. The diff appends a new perf-changelog entry whose description list contains both an English line and a Chinese translation line, under the PR 3401 entry in perf-changelog.yaml. AGENTS.md:126 states new entries must be English-only, with no Chinese translations or bilingual descriptions. Neither validate_perf_changelog.py nor validation.py contains any bilingual check, so the line causes no validation failure. Fix is to remove the Chinese description string, keeping only the English line.
|
/reuse-sweep-run 36538876933 |
functionstackx
left a comment
There was a problem hiding this comment.
why nixl patch and why so many chnages outside of draft preicison
Description
Preserve GLM-5.2 GB200 draft precision and select the guarded UCX prefill workaround for the same two disaggregated recipes through main's Python launcher.
Testing: Sweep 36538876933/a3, head
9fbe9bf0: 14 performance points and four evals passed with per-cell artifact/runtime review. Current submission regressions and changelog/matrix validation pass.Review limits: The engine-patch waiver remains pending. Consolidation contains old/retried C8 rows; select final job
110329500299before ingestion. Identity, power, cancellation and shutdown limits remain. The new launcher/runtime and inherited telemetry changes lack fresh GPU qualification.中文
保留 GLM-5.2 GB200 draft 的原始发布精度,通过 main 的 Python launcher 为原有两个分离式配方选择受哈希保护的 UCX prefill 规避脚本。
测试: sweep
36538876933/a3(9fbe9bf0)的 14 个性能点和 4 个 eval 通过,已逐格审查 artifact/runtime。当前提交参数回归及 changelog/矩阵验证通过。审阅限制: engine-patch 例外仍待批准。合并结果包含 C8 原始/重跑记录,入库前须选择最终 job
110329500299。身份比对、功耗、取消及关闭限制保留;新 launcher/runtime 和继承的遥测变更尚无新 GPU 验证。AI 模型: Claude Opus 5.5 (
claude-opus-5-5) 负责实施及起草;GPT-6(具体变体不可确认)负责恢复、验证、冲突解决及委派复核。AI model disclosure
claude-opus-5-5): implementation/drafting.Related Issue
Related to #3228. / 关联 #3228。
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.